Papers with speech enhancement

4 papers
Speaking in Wavelet Domain: A Simple and Efficient Approach to Speed up Speech Diffusion Model (2024.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to enhance inference speed and training require complex modifications to the model.
Approach: They propose to double the training and inference speed of Denoising Diffusion Probabilistic Models by simply redirecting the generative target to the wavelet domain.
Outcome: The proposed method doubles the training and inference speed of Speech DDPMs by redirecting the generative target to the wavelet domain.
Far-Field Speaker Recognition Benchmark Derived From The DiPCo Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Using a publicly-available corpus, we propose a far-field speaker verification benchmark.
Approach: They propose a far-field speaker verification benchmark derived from the publicly available DiPCo corpus.
Outcome: The proposed tasks are very challenging and hope to inspire the speech community to develop new methods and systems for this challenging domain.
SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing (2022.acl-long)

Copied to clipboard

Challenge: Existing work shows that pre-trained models can improve in various natural language processing tasks.
Approach: They propose a unified-modal encoder-decoder framework that pre-trains speech-text representations using large-scale unlabeled speech and text data.
Outcome: The proposed framework is superior to existing models on speech-to-text processing tasks.
LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech Enhancement (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in language models have demonstrated strong capabilities in semantic understanding and contextual modeling.
Approach: They propose a LLaMA-based language model that incentivizes generalization capabilities for speech enhancement.
Outcome: The proposed language model outperforms prior task-specific discriminative and generative models in acoustic enhancement tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations